Alibaba Launches Qwen3.8-Omni-Flash: Native Multimodal, Million-Context Audio Cost Cut by 98%
The Alibaba Qwen team has launched the native multimodal model Qwen3.8-Omni-Flash, which supports four types of input: text, image, audio, and video, with a context window of up to 1 million Tokens. It achieves an average improvement of over 25% across 29 benchmarks compared to its predecessor. The most significant change lies in its processing approach: instead of simply concatenating speech and image recognition, it natively integrates multimodal understanding and AI agent capabilities at the model level.